.png)



Explore the world of local LLM inference and understand what happens when AI models run on your own hardware. We’ll explore model architectures, checkpoints, formats such as GGUF, MLX, Safetensors, GPTQ and AWQ, and runtimes like LM Studio and Ollama. The session covers inference mechanics, prefill, decode, KV cache, context, sampling, latency, throughput, and dense vs MoE models. Through hands-on experiments, we’ll connect these concepts to Apple Silicon, GPUs, VRAM, unified memory, bandwidth, and real-world inference performance.